Back

Journal of Speech, Language, and Hearing Research

American Speech Language Hearing Association

Preprints posted in the last 7 days, ranked by how well they match Journal of Speech, Language, and Hearing Research's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Evaluating Goodness of Pronunciation and Phonological Posteriors as Objective Markers of Speech Severity in Motor Speech Disorders

Wang, F.; Utianski, R. L.; Duffy, J. R.; Barnard, L. R.; Botha, H.

2026-07-16 neurology 10.64898/2026.07.14.26358076 medRxiv
Top 0.1%
23.3%
Show abstract

This study examined the extent to which goodness of pronunciation (GoP) scores and phonological posterior probabilities capture perceptual ratings of speech severity in individuals with motor speech disorders (MSD). Speech recordings of the word catastrophe were obtained from 489 participants, including 333 neurologically typical controls and 156 individuals with MSD. GoP scores were derived using traditional acoustic features and self-supervised speech representations, including WavLM and XLS-R, across multiple modeling approaches, while phonological posterior probabilities were extracted using Phonet. Model performance was evaluated using Kendall's rank correlations, regression, and receiver operating characteristic analyses against speech-language pathologists' perceptual ratings of sound distortion and intelligibility. Both GoP and phonological posterior probabilities were significantly associated with perceptual ratings. Self-supervised speech representations substantially outperformed traditional acoustic features, with WavLM-based GoP using k-nearest neighbors achieving the strongest performance. Across correlation, regression, and classification analyses, GoP consistently outperformed phonological posterior probabilities for both sound distortion and intelligibility. Age and gender had minimal influence on model-derived measures or their relationships with perceptual ratings. These findings demonstrate the value of self-supervised GoP as an objective measure of speech impairment while highlighting the complementary role of phonological posterior probabilities in characterizing articulatory aspects of motor speech disorders.

2
Automated Detection of Motor Speech Disorders and Subtype Classification

Wang, F.; Utianski, R. L.; Barnard, L. R.; Stricker, J. L.; Clark, H. M.; Meade, G. F.; Jones, D. T.; Whitwell, J. L.; Josephs, K. A.; Duffy, J. R.; Botha, H.

2026-07-19 neurology 10.64898/2026.07.16.26358268 medRxiv
Top 0.1%
7.8%
Show abstract

Motor speech disorders (MSDs) are early markers of neurological disease, but expert perceptual analysis is rarely available outside specialized centers. Automated speech analysis offers a scalable alternative, yet prior studies have not systematically compared modeling approaches or assessed clinically relevant metrics in independent datasets. This study compared static acoustic features, articulatory informed Phonet features, and self-supervised pretrained models for binary and multi label MSD classification. We trained and evaluated models on 583 speech samples using speaker level splits. Baseline models included logistic regression and Gated Recurrent Units (GRUs) trained on eGeMAPS and MFCCs. We extracted three types of Phonet derived features and evaluated pretrained HuBERT and SSAST models in frozen, partially fine-tuned, and fully fine-tuned configurations. Binary classification distinguished MSDs from controls, while multi label classification identified six MSD subtypes. Models were assessed using validation AUC, and cut points were tested on two independent datasets. Pretrained and Phonet based models substantially outperformed static acoustic features. In binary classification, HuBERT achieved the highest AUC (0.95), while compact Phonet derived GRUs achieved comparable performance (up to 0.94). These models generalized well to independent datasets, maintaining high sensitivity (0.94) and specificity (0.97). In multi label classification, Phonet models achieved the highest macro average AUC (0.86), but threshold-based subtype performance declined on unseen data. Automated MSD detection is feasible and clinically promising. Binary classification generalized well, whereas multi label classification showed limited threshold stability across datasets.

3
Validity and Reliability of the Novel Indonesian Instrument for Aphasia Diagnosis (IDEA)

Prawiroharjo, P.; Fakhri, A.; Gabrielle, A.; Martalia, V.; Rahmayani, S. A.; Wijaya, V. G.

2026-07-19 neurology 10.64898/2026.07.17.26358303 medRxiv
Top 0.1%
2.9%
Show abstract

Aphasia diagnosis in Indonesia remains challenging due to limited culturally and linguistically appropriate instruments. Widely used tools such as the Boston Diagnostic Aphasia Examination (BDAE) and Western Aphasia Battery (WAB) are not adapted to the Indonesian context, while Tes Afasia untuk Diagnosis, Informasi, dan Rehabilitasi (TADIR) provides screening but lacks diagnostic accuracy. To address this gap, we developed the Instrumen Diagnosis dan Evaluasi Afasia (IDEA) for native Indonesian speakers and evaluated its validity, reliability, and normative cutoff values in cognitively healthy Indonesian adults. Eighty-three cognitively normal adults (screened using MoCA-Ina) with no history of neurological disease were assessed using IDEA, which evaluates six language domains. Items were adapted from existing tools and reviewed by experts. Content validity, internal consistency (Cronbachs alpha), and construct validity (Exploratory Factor Analysis) were analyzed using SPSS v25. A total of 83 participants were included (median age = 55.81 years, 54% secondary education). IDEA demonstrated good feasibility, with an average completion time of 45-60 minutes depending on participant engagement. Content validity was established by unanimous expert consensus. Construct validity showed meritorious sampling adequacy (KMO = .872) and significant sphericity (Bartletts test {chi}^2 (15) = 278.523, p<.001), supporting factor analysis. Internal consistency showed good reliability across six domains (Cronbachs = 0.896). IDEA is a valid and reliable tool for assessing aphasia in Indonesian natives. It is a culturally appropriate assessment tool which offers structured, domain-based evaluation and supports differential diagnosis of both classical and progressive aphasia syndromes. Keywords: Aphasia, Language Assessment, Indonesian, IDEA, Validity

4
Benchmarking Speech Recognition Models for Medical Consultations in Latin American Spanish: A Comparative Evaluation with Fine-Tuning

Carrillo, R. M.; Carbajal Serrano, A.; Condori Pinedo, P. S.

2026-07-16 public and global health 10.64898/2026.07.14.26358062 medRxiv
Top 0.1%
1.2%
Show abstract

BACKGROUND: Artificial intelligence (AI) medical scribes rely on speech-to-text (STT) models for transcription. Evaluations of STT models in non-English settings remain scarce. We benchmarked ten STT models on medical consultations from Latin American (LatAm) Spanish and assessed whether fine-tuning improves transcription accuracy. METHODS: Ten YouTube videos depicting medical consultations. Human transcriptions were the ground truth. Five open-source models were evaluated: Whisper Large, Whisper Large v3, Whisper Large v3 Turbo, Voxtral Mini 3B, and Canary 1B v2; and so were five close-source models: gpt-4o-transcribe, gpt-4o-mini-transcribe, gemini-2.5-pro, Eleven Labs, and Assembly AI. Whisper Large v3 was fine-tuned. One video was withheld from training. Performance assessed using Word Error Rate (WER), Character Error Rate (CER), BLEU Score, ROUGE-L, BERT Score, and Semantic Similarity on the one withheld video. RESULTS: None of the fine-tuning iterations outperformed the vanilla Whisper Large v3. With the withheld video, Gemini-2.5-pro was the close-source model with the best performance in four of six metrics. In comparison to the close-source models, the fine-tuned model never outperformed the other models (withheld video); conversely, in comparison to the close-source models, the fine-tuned model showed better performance across metrics, for instance: BLEU score (63% vs to 58% for the second-ranking model), BERT (89% vs to 86%), and semantic similarity (89% vs to 83%), CER (19% vs 20%). CONCLUSIONS: Whisper Large v3 and its fine-tuned variant are the best open-source STT models for transcribing medical conversations in LatAm Spanish. These findings provide an evidence base for developing AI medical scribes tailored to Spanish-speaking LatAm.

5
Autism Research at a Crossroads: Global Progress, Persistent Gaps, and Future Pathways: A Bibliometric Analysis

zhong, Q.; Chen, L.; Ji, Y.; Zhu, F.; Zou, X.

2026-07-16 psychiatry and clinical psychology 10.64898/2026.07.14.26358066 medRxiv
Top 0.3%
0.5%
Show abstract

Background The global prevalence of autism spectrum disorder (ASD) has significantly increased over the past two decades. Despite substantial research advances, critical aspects, including etiology, diagnostic biomarkers, and pharmacological interventions, remain incompletely elucidated. This persistent knowledge gap warrants systematic mapping of the field's evolution to inform future research priorities. Methods A bibliometric analysis of ASD-related publications indexed in Web of Science was conducted from January 2020 to May 2025. Following a systematic deduplication process, original articles, reviews, case reports, and clinical trials were included in the analysis. The analytical framework comprised co-authorship networks, institutional collaboration patterns, national research contributions, and keyword co-occurrence structures, all of which were examined using CiteSpace (version 5.8.R3) and VOSviewer. Results After deduplication, 8,162 publications (January 2020-May 2025) were analyzed. The annual output grew steadily, confirming ASD as a sustained priority in neuroscience. Research remains academia-driven, led by the United States, with China as the second-largest contributor. Chinese institutions place greater emphasis on mechanistic and developmental phenotyping, which aligns with national priorities. These studies maintain strong methodological rigor, and their growing volume underscores the central role of ASD in translational neuroscience. Conclusion Future research on ASD should focus on strengthening case identification, refining clinical phenotyping, and expanding large-scale cohort studies to advance our understanding of its etiology and identify reliable diagnostic biomarkers. It is equally important to develop and evaluate targeted interventions for core symptoms and integrate telemedicine into service delivery models. A critical yet understudied priority is improving the quality of life for autistic individuals and their families, an area in which research globally, including in China, requires greater depth and consistency. With China's growing investment in autism research, it is well-positioned to contribute to these pressing international challenges.

6
The Shape of a Final Message: An Emotional Landscape in the Language of Suicide

Pestian, J. P.; Jacobson, D. A.; Pedapati, E. V.; Mendonca, E. A.; McMahon, B. H.; Ive, J.; Glauser, T. A.

2026-07-17 psychiatry and clinical psychology 10.64898/2026.07.16.26358230 medRxiv
Top 0.5%
0.2%
Show abstract

The emotional content of suicide notes is typically examined using categorical coding, where each labeled passage is treated in isolation from its surrounding language. In contrast, dimensional models of psychopathology propose that affective content varies along continuous gradients. We evaluated this proposition directly. Excerpts from 884 annotated suicide notes were embedded in a semantic space defined solely by their linguistic properties, and we investigated whether human-assigned emotion labels changed smoothly across this space. They did: affective tone showed clear spatial autocorrelation (Moran's $I = 0.18$, $z = 19.68$, $p < 0.001$), an effect that replicated across three different encoders and remained after removing all within-note dependencies. Emotions occupied recognizable yet overlapping regions rather than forming distinct clusters and varied substantially in how tightly they were concentrated: love and hopelessness appeared with similar frequency, but love was far more localized ($z = 15.7$ versus $10.8$). Among all emotions, hopelessness was the most linguistically diffuse, implying that a single categorical label is capturing multiple, qualitatively different manifestations of suicidal distress.

7
Are CNV Risk Scores Linked to Neurodevelopmental and Mental Health Characteristics Within CNV-Associated Intellectual Disability?

Chi, Z.; Alexander-Bloch, A.; Neufeld, S. A.; Wolstencroft, J.; Skuse, D.; IMAGINE-ID consortium, ; Baker, K.

2026-07-16 psychiatry and clinical psychology 10.64898/2026.07.14.26358034 medRxiv
Top 0.6%
0.1%
Show abstract

Background: Children and young people (CYP) with intellectual disability (ID) frequently have co-occurring neurodevelopmental (ND) and mental health (MH) difficulties. While copy number variants (CNVs) are identified as an important aetiology of ID, it is unclear whether and how CNV risk scores predict ND and MH characteristics within the CNV-associated ID population. Methods: We analysed data from the UK-based IMAGINE-ID cohort of CYP (aged 4-19 years) with ID and clinically-reported CNVs (N = 1,640). CNVs were annotated with Gencode 19 in ENSEMBL to calculate CNV risk scores, including summed probability of loss-of-function intolerance (pLI) and dosage sensitivity. Multivariate regression models examined the prediction of CNV variables and inheritance on ND and MH characteristics, assessed via the Development and Well-Being Assessment (DAWBA). Post-hoc analyses explored CNV variable stratification (lower vs. higher range pLI). Results: Higher summed pLI scores (indexing CNV genes' intolerance to loss of function) unexpectedly predicted fewer MH difficulties and a lower likelihood of ND diagnoses, even after accounting for demographic factors and CNV inheritance. Post-hoc analyses identified a threshold effect. Within the lower pLI range, higher pLI scores were associated with greater MH difficulties, consistent with findings from population-based samples. In contrast, within the higher pLI range, higher pLI scores were associated with fewer MH difficulties (among individuals more likely to have severe ID). Conclusion: These findings challenge the assumption that CNV genomic "risk scores" universally predict ND and MH difficulties. Instead, within CNV-associated ID, complex relationships exist between CNV risk scores, inheritance and phenotypes. These insights emphasise the necessity of integrating genomic results with familial and developmental context to understand individual vulnerabilities and support needs.

8
Machine learning and data-driven models for predicting post-stroke dysphagia: a systematic review and meta-analysis

Mohammadi Yazdi, S.; Motevaselian, M.; Khatami, S.; Radfar, N.; jourahmad, z.; Perez, H. A.

2026-07-17 neurology 10.64898/2026.07.15.26358113 medRxiv
Top 0.7%
0.1%
Show abstract

Background: Post-stroke dysphagia (PSD) contributes to aspiration, pneumonia, malnutrition, prolonged hospitalization and mortality. We evaluated the discrimination, validity and readiness of machine learning and data-driven prediction models for PSD-related outcomes. Methods: Following a prospectively registered protocol (PROSPERO CRD420261419259), we searched PubMed/MEDLINE, Embase, Web of Science Core Collection, CINAHL and CENTRAL from inception through June 7, 2026. Eligible studies developed or validated multivariable prediction models for PSD-related outcomes in adults with stroke. We used PROBAST and PROBAST+AI to assess risk of bias and applicability and TRIPOD+AI to evaluate reporting. Area under the curve (AUC) estimates were pooled on the logit scale with random-effects models. Results: Twenty-four studies were included and ten contributed to meta-analysis. Four studies predicting early or incident PSD yielded a pooled AUC of 0.94 (95% CI 0.60-0.99; I2 = 95.6%). Pooled AUCs were 0.84 (95% CI 0.71-0.92) for aspiration or penetration-aspiration and 0.89 (95% CI 0.24-1.00) for severe dysphagia. The exploratory analysis of all ten risk-prediction models produced an AUC of 0.90 (95% CI 0.80-0.95), but heterogeneity was substantial (I2 = 90.3%) and the prediction interval was 0.51-0.99. Every study had high risk of bias because of analysis-domain concerns; calibration and external validation were uncommon. Conclusions: Reported discrimination was often high, but the evidence does not establish reliable performance in care. Independent validation, calibration, complete model reporting and clinical-impact studies are needed before these models guide post-stroke swallowing care. Keywords: Post-stroke dysphagia; Stroke; Deglutition disorders; Machine learning; Clinical prediction model; Area under the curve; Meta-analysis

9
Prompt Engineering Limitations: Preliminary Evaluation of Large Language Models for Psychotherapy Safety

Ngo, N.; Dao, G.; Sano, A.

2026-07-18 psychiatry and clinical psychology 10.64898/2026.07.16.26358261 medRxiv
Top 0.9%
0.1%
Show abstract

Large Language Models are increasingly used in consumer-facing mental health tools, many of which claim that prompt engineering alone can ensure safe therapeutic behavior. This study evaluates that assumption by testing 20 proprietary and open-source LLMs on high-risk psychiatric scenarios, using prompts grounded in behavioral therapy principles. Prompt engineering reduced some predictable risks, such as explicit endorsement of self-harm, but consistently failed in ambiguous or clinically nuanced situations. Models frequently validated harmful statements, colluded with hallucinations, minimized symptoms, or used stigmatizing language, including in the newest and largest models. These failures reflect structural limitations such as lack of memory, insufficient contextual reasoning, and training-related biases. Prompt engineering alone is therefore insufficient for safe AI-mediated psychotherapy; clinician-guided fine-tuning, integrated safety mechanisms, and system-level oversight will be required. This work provides early evidence motivating deeper clinician-led evaluation and safety-oriented model development.

10
A Culturally Embedded Augmented Reality Task as a Neurocognitive Biomarker of Executive Function in Schizophrenia

Chatthong, W.; Rueankam, M.; Khemthong, S.

2026-07-16 psychiatry and clinical psychology 10.64898/2026.07.14.26358053 medRxiv
Top 1.0%
0.1%
Show abstract

Executive function (EF) deficits are central features of schizophrenia and strongly influence long-term functional outcomes. Conventional cognitive assessments often lack ecological validity and cultural relevance. This study introduces the Luk Chup Augmented Reality (LCAR) tool a video guided, clay modeling task delivered through wearable AR that integrates culturally familiar activity with realtime neurophysiological monitoring. Thirty individuals diagnosed with schizophrenia (mean age = 38.9, SD. = 7.15 years) completed a series of modeling and memory tasks using LCAR while undergoing quantitative EEG (QEEG). Task duration and theta/beta power were analyzed across procedural and color shape memory phases. Memory phases took significantly longer to complete and were associated with decreased lateral prefrontal theta and increased frontal midline theta activity (Fz, Cz), indicating higher EF demand. A repeated-measures ANOVA revealed significant condition, site, and interaction effects on theta power. The LCAR tool shows promise as a culturally grounded, dual-mode assessment of EF in schizophrenia. It offers a novel integration of performance-based and neurophysiological metrics that may inform future interventions in psychiatric rehabilitation.

11
Efficient stochastic epidemic simulation via the Sellke construction

van Boven, M.; Bootsma, M. C.

2026-07-17 epidemiology 10.64898/2026.07.16.26358219 medRxiv
Top 1%
0.0%
Show abstract

Stochastic epidemic models are a cornerstone of infectious disease epidemiology and are often used to study intervention scenarios. However, large run-to-run variability can make intervention effects difficult to estimate precisely. We revisit the epidemic Sellke construction, which assigns each individual an infection threshold for the cumulative infection hazard such that, conditional on the thresholds, the epidemic trajectory becomes deterministic. This enables coupling of simulations with and without an intervention, yielding low-variance effect estimates even when outcomes such as final size or peak incidence vary widely between runs. We develop an exact, event-driven implementation that maintains infection and recovery events in priority queues. Cumulative infection-hazard updates require O(log N) time per event, yielding overall complexity O(Elog N) for E events in a population of size N. The implementation achieves computational performance comparable to the classical Gillespie algorithm while naturally accommodating non-Markovian infectious periods and complex infectiousness profiles. We illustrate the approach using distance-dependent spread of avian influenza between poultry farms in the Netherlands and a multilayer population with households, schools, and workplaces. In both examples, coupling enables efficient within-run comparisons of intervention scenarios across stochastic realisations.

12
Comparative Efficacy of Vancomycin and Fidaxomicin Regimens for the Prevention of Recurrent Clostridioides difficile Infection: A Systematic Review and Network Meta-Analysis of Randomized Controlled Trials

Prosty, C.; Butler-Laporte, G.; Brophy, J.; Frenette, C.; Loo, V.; Coburn, B.; Hota, S.; Longtin, Y.; Kong, L.; Muller, M.; Steiner, T.; Valiquette, L.; Daneman, N.; Daley, P.; Nott, C.; MacFadden, D. R.; Kandel, C.; Chen, Y.; Perez- Patrigeon, S.; Lee, T. C.; McDonald, E.

2026-07-17 infectious diseases 10.64898/2026.07.14.26358112 medRxiv
Top 1%
0.0%
Show abstract

Background and Aims The optimal treatment for first episodes and first recurrences of Clostridioides difficile infections (CDI) is unknown and there is emerging evidence for pulse and taper (P-T) regimens. Therefore, we sought to estimate the relative efficacy of treatment options. Methods MEDLINE and CENTRAL were searched from database inception to May 21, 2025 and unpublished conference abstracts were searched from recent infectious disease conferences. RCTs on the treatment of first episodes or first recurrences of CDI comparing fixed-dose or P-T regimens of fidaxomicin or vancomycin were included. The primary and secondary outcomes were 40- and 56-day CDI recurrence, respectively. A random-effects network meta-analysis on the risk ratio (RR) scale was conducted using a standard regimen (10-14 days) of vancomycin as the comparator. Treatments were ranked using the surface under the cumulative ranking curve (SUCRA). Results 8 RCTs were included comprising a total of 2181 patients. For 40-day recurrence, fidaxomicin P-T had the highest probability of ranking best (RR=0.10, 95%Confidence Interval [95%CI]=0.10-0.49, SUCRA=1.00), followed by vancomycin P-T (RR=0.49, 95%CI=0.32-0.76, SUCRA=0.61), fixed-dose fidaxomicin (RR=0.61, 95%CI=0.49-0.76, SUCRA=0.39), and, finally, fixed-dose of vancomycin (SUCRA=0.00). The treatments ranked in the same order for 56-day recurrence, though only 3 RCTs reported on this timepoint. Conclusion Vancomycin P-T, fidaxomicin P-T, and fixed-dose fidaxomicin were all superior to a fixed-dose vancomycin. Head-to-head comparative effectiveness RCTs are needed to quantify their relative effect sizes of and impact on long-term prevention of recurrent CDI.

13
Nationwide Mpox Genomic Surveillance Reveals Clade Ib Introductions, APOBEC3-Driven Evolution, and Terminal Deletions

Brochu, H. N.; Shi, Q.; Song, K.; Zhang, Q.; Munroe, J.; Harris, N. J.; Britt, N.; Zeng, Q.; Kapuria, K.; Chappell, J.; Norvell, B. M.; Peavy, L.; Williams, J. D.; Harris, A. B.; Chaitram, J.; Hutson, C. L.; Deng, J.; McGrath, D.; Boles, D.; Dale, S. E.; Gigante, C. M.; Iyer, L. K.

2026-07-17 infectious diseases 10.64898/2026.07.15.26357894 medRxiv
Top 1%
0.0%
Show abstract

Background The 2022-2023 global mpox outbreak highlighted the critical need for robust genomic surveillance capabilities to track mpox virus (MPXV) evolution and transmission dynamics. Methods Building upon our established SARS-CoV-2 sequencing infrastructure, we implemented a Molecular Loop probe-based long-read sequencing approach using Pacific Biosciences Sequel II technology for comprehensive MPXV genomic surveillance across the United States (US). From August 2024 to June 2025, we generated 326 high-quality whole genome sequences from residual mpox-positive clinical specimens collected by Labcorp across all 10 US Department of Health and Human Services regions. Results Our analysis identified two samples containing clade Ib MPXV in January and June 2025 and captured shifting trends in clade IIb diversity, with 13 distinct lineages observed. We also identified multiple instances of large (~1.6-17.6kb) deletions proximal to the inverted terminal repeats in clade IIb genomes. APOBEC3 mutation analysis indicated substantial evidence of human-to-human transmission among both clades. Further, we observed significantly higher APOBEC3-associated SNPs per kilobase (P<0.001) in clade IIb genomic variable regions relative to their central conserved region. Our assay exhibited strong reproducibility across biological replicates from individual patients and accuracy was confirmed via parallel sequencing of select specimens by US Centers for Disease Control and Prevention (CDC) using metagenomic sequencing. We also demonstrated via custom simulation that our assay discriminates all known MPXV clades and lineages, including those we have not observed in the US. Conclusions Our integrated nationwide surveillance system facilitates real-time genomic tracking of outbreak evolution, with demonstrated capacity across SARS-CoV-2 and MPXV, positioning this platform for rapid deployment during future pathogen emergence.

14
Complex intra-host SARS-CoV-2 evolution following monoclonal antibody pre-exposure prophylaxis

Kamelian, K.; Pascall, D. J.; Cheng, M. T. K.; Meng, B.; Altaf, M.; Morse, R. M.; Aggio, J. B.; Egan, D. J. S.; Chen-Xu, M.; Trivioli, G.; Sutton, B.; Richter, A.; Gonzalez-Vazquez, L. D.; Cormie, C.; Kemp, S.; Yeadon, R.; Hyatt, B.; Wong, A.; Thesin Pelamkulangara, N.; Fraser, E.; McCarthy, B.; Novaes, F.; Stott, S.; Galvin, A.; Bellis, K. L.; De Angelis, D.; Harrison, E. M.; Martin, D.; Smith, R. M.; Gupta, R. K.

2026-07-17 infectious diseases 10.64898/2026.07.14.26356329 medRxiv
Top 1%
0.0%
Show abstract

Background: Monoclonal antibodies have emerged as a prophylactic strategy to prevent symptomatic SARS-CoV-2 infection in immunocompromised individuals. However, the evolutionary and clinical implications of breakthrough infections under this regime remain unclear. Methods: A male in their 80s with a haematological/oncological diagnosis received a 2000 mg intravenous infusion of sotrovimab in March 2023 and was diagnosed with COVID-19 by RT-qPCR from a nasopharyngeal swab in August 2023. Weekly samples (n=24) were collected through February 2024 (171 days). All samples underwent whole-genome sequencing, with select mutations subjected to functional assessment. Findings: Sequencing identified the GE.1 lineage at all timepoints. An intra-host recombination event in ORF1ab (positions 8942-12458) was detected prior to 23 weeks post-detection, followed by a 14-fold increase in viral load (7.42e+06 to 1.00e+08 RNA copies/mL) and a marked shift in the viral population. E340D, a sotrovimab resistance mutation, was detected at low abundance (46%) within the first week post-infection, fluctuated over time, and was nearly fixed by week 15 (107 days) post-detection. We assessed five spike mutations - V36M, S98F, and V213G in the N-terminal domain, Y505P in the receptor-binding domain, and P681Q near the S1/S2 cleavage site - and additionally evaluated the impact of E340D. V36M conferred the highest infectivity across all cell lines, with the most significant effect in low-TMPRSS2 cells. While all mutations showed enhanced infectivity with the addition of E340D, the effect was most pronounced in mutations with lower baseline infectivity. The addition of E340D significantly decreased relative neutralizing titres for V36M, S98F, and V213G, enabling escape from neutralizing antibodies in XBB-responsive individuals, illustrating an enhanced phenotypic advantage. Patient neutralizing activity was absent pre-sotrovimab, and sotrovimab-induced neutralization was further compromised by selection of E340D. Interpretation: Sotrovimab pre-exposure prophylaxis in an immunocompromised patient did not prevent SARS-CoV-2 infection, and selected for resistant mutation E340D, with unexpected fitness consequences across non-receptor binding domain spike regions.

15
Genome-Wide Association Studies and Deep-Learning Functional Annotation of Opioid Use Disorder across Three Ancestries in the All of Us Research Program

Gu, S.; Petrovitch, D.; Hall, O. T.; Lambert, J. W.; Kember, R. L.; Nahid, N. A.; Ma, Q.; Sprague, J. E.; McDonough, C. W.; Johnson, J. A.

2026-07-17 addiction medicine 10.64898/2026.07.15.26358096 medRxiv
Top 1%
0.0%
Show abstract

Background: Opioid use disorder (OUD) is heritable, yet most genome-wide association studies (GWAS) have focused on European populations, leaving the genetic architecture of OUD in non-European populations underexplored. Methods: We conducted GWAS of OUD across three ancestries using electronic health records and genomic data from 52,357 All of Us Research Program participants (8,912 cases; 43,445 matched opioid-exposed controls; 48.5% female). Participants were stratified into European (EUR), African (AFR), and Admixed American (AMR) ancestry groups for logistic regression GWAS, with independent replication in the Million Veteran Program. We then applied the deep-learning model AlphaGenome to predict the tissue-specific transcriptomic and splicing consequences of top risk variants across 13 reward-pathway brain regions. Results: We identified and replicated a novel DDX6 risk locus, alongside established OPRM1 and FURIN signals. AlphaGenome predicted the DDX6 regulatory allele downregulates the stress-resistance gene FOXR1 in the nucleus accumbens, while the protective OPRM1 variant (rs1799971) upregulates OPRM1 expression across reward networks. Other signals of interest included IL6R and SHISA9 (EUR); GHR (AFR); and ASTN2 (AMR). Conclusions: This study identifies DDX6 as a novel OUD risk locus, replicates associations with OPRM1 and FURIN, and highlights biologically plausible ancestry-specific signals in AFR and AMR populations. We also replicated top variants in an independent population. Finally, integrating GWAS with deep-learning annotations provides specific, localized biological hypotheses to guide future experimental validation and targeted therapeutics.

16
Molecular and phylogenetic insights into the novel Brugia sp. in Sri Lanka with new evidence for zoonotic transmission

Nimalrathna, S. U.; Harischandra, H.; Kimber, M.; Chandrasena, N.; De Silva, N.; Mallawarachchi, H.; De Silva, B. G. D. N. K.

2026-07-21 infectious diseases 10.64898/2026.07.20.26358473 medRxiv
Top 1%
0.0%
Show abstract

The World Health Organization (WHO) validated Sri Lanka had eliminated lymphatic filariasis as a public health problem in 2016, the second country in Southeast Asia to attain this status. However, post-validation surveillance has identified sporadic cases of brugian filariasis. The reemergence of Brugia malayi infections in Sri Lanka warrants urgent investigations. Recent studies have shown that the parasite responsible for the reemergence is a novel zoonotic Brugia sp. maintained among dogs that is closely related but distinct to the human-infecting B. malayi species. The current study employed morphological and morphometric assessments, revealing that this novel zoonotic Brugia sp. is within the B. malayi morphological range. Molecular characterization of three genomic regions, the nuclear genomic region SLXI, the non-coding region HhaI, and the mitochondrial genomic region COXI confirmed it as a genetic variant more closely related to B. malayi than to B. pahangi. Phylogenetic analysis further indicated it as a distinct genomic variant, closely related to a B. malayi-like parasite reported from India. Notably, that same parasite was identified in infected humans, animals, and potential vector mosquitoes. This, together with the detection of both human and animal blood within the same brugian infective mosquitoes, and delineating the canine origin of the parasites in human infections, provides compelling evidence supporting zoonotic transmission of this parasite. To our knowledge, this is the first report demonstrating the presence of the same brugian parasite in humans, domestic animals, and potentially infective mosquitoes in Sri Lanka, supported by multi-genomic evidence. The recent identification of multiple potential mosquito vector species suggests that this parasite may have undergone adaptive changes, facilitating its ability to overcome the species barrier. These findings substantiate the long-held hypothesis of zoonotic transmission of the reemerged brugian parasite, highlighting significant implications for ongoing surveillance and control strategies.

17
Bridging surveillance gaps in dengue: a hierarchical model integrating mixed data sources for transmission estimation and vaccine targeting

Djaafara, B. A.; Elyazar, I. R.; Yosephine, P.; Surya, A.; Silalahi, F. S.; Handito, A.; Thohir, B.; Aryani, D.; Gunawan, D.; Nisa, A. K.; Prianto, E.; Samad, I.; Cook, A. R.; Huang, A. T.; Clapham, H. E.; Bhatt, S.; Mishra, S.

2026-07-17 epidemiology 10.64898/2026.07.15.26358208 medRxiv
Top 1%
0.0%
Show abstract

Estimating dengue force of infection (FOI) is essential for understanding transmission dynamics and targeting intervention programmes, yet surveillance data in endemic settings required for estimations are often incomplete, with varying formats. We developed a Bayesian hierarchical catalytic model that jointly fits age-stratified case data, aggregate case data, and seroprevalence surveys within a single framework, incorporating external covariates to improve parameter identifiability. Synthetic validation showed that covariates alone recovered accurate FOI point estimates even when most districts contributed only aggregate data, but did so with poorly calibrated uncertainty; anchoring the model with a single seroprevalence survey was necessary to bring credible interval coverage close to nominal. Applied to 128 districts across Java and Bali, Indonesia (2016-2024), the model revealed substantial spatial heterogeneity in FOI and reporting rates. Many districts in Java exceeded the WHO-suggested seroprevalence threshold for vaccine introduction, yet were classified as low-priority when using reported incidence as prioritisation criterion, particularly in areas with weak surveillance. Model-based seroprevalence estimation, integrating multiple data sources, offers a more consistent basis for identifying high-priority districts for vaccine introduction, and is less susceptible to surveillance bias than reported incidence.

18
Rationale and guidance for implementing the continual reassessment method for dose-finding in controlled human infection model studies

Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.

2026-07-17 infectious diseases 10.64898/2026.07.16.26358128 medRxiv
Top 1%
0.0%
Show abstract

Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.

19
General Practice Perspectives on Post-Infection Conditions: Scoping Review and UK Survey

Aung, K. W.; Scuffell, J.; Podlasek, A.; Engamba, S.; Jones, F.; Edwards, A.; Chew-Graham, C. A.; Sanyaolu, L.; Busse-Morris, M.

2026-07-17 primary care research 10.64898/2026.07.15.26358157 medRxiv
Top 1%
0.0%
Show abstract

Background Post-infection conditions (PICs), such as Long Covid, are associated with heterogeneous, fluctuating symptoms that profoundly affect daily functioning. Despite moderate-certainty evidence from the NIHR-funded LISTEN trial (COV-LT2-0009) that personalised self management support improves outcomes and may reduce societal and economic impacts of Long Covid, many people living with PICs still receive condition-specific services, generic advice, or stand-alone digital tools that do not address their complex needs. Aim To map care approaches in general practice and synthesise UK evidence for PIC management. Design and setting Scoping review and online survey. Method A two-phase study was conducted: (1) a scoping review of UK evidence on PIC management in general practice; and (2) a supplementary online survey of practitioners working in UK general practice to provide contextual insights. Results The scoping review identified 32 studies focused on Long Covid. One study included a comparator group (ME/CFS). Study populations were predominantly white ethnicity and female. Evidence for non-Covid PICs in UK general practice was largely absent. The supplementary survey (n=46) provided preliminary practice-level insights. Healthcare practitioners reported varied PIC presentations, diagnostic uncertainty, limited referral pathways, inequitable access, and low confidence in managing PICs. Conclusion Evidence informing PIC management in UK general practice remains predominantly Long Covid-focused and may not reflect the range of PICs encountered in practice. While survey findings are preliminary and require confirmation in larger samples, they highlight uncertainty around PIC management. Further research is needed to evaluate whether existing Long Covid pathways should be expanded or complemented by broader PIC models. Keywords general practice; Long Covid; self-management; post-viral syndromes

20
Temporal relationships between distress and pain in people living with HIV

Arendse, G.; Kamerman, P.; Wadley, A.; Edwards, R. R.; Joska, J.; Parker, R.; Madden, V. J.

2026-07-17 primary care research 10.64898/2026.07.15.26358133 medRxiv
Top 1%
0.0%
Show abstract

Objective: There is a bidirectional relationship between emotional distress and pain. However, this relationship is understudied in people with HIV in low-resource settings. This study sought to describe the temporal relationship between emotional distress and pain in people with HIV. Design: Longitudinal observational study. Methods: Participants with virally suppressed HIV, reporting either no pain or persistent pain at baseline, provided weekly remote ratings of distress, worst pain, and average pain using 0-10 visual analogue scales. Within-individual fluctuations in distress and pain were visualised over time. Group-level correlations were determined using Spearman's correlation tests. Cumulative link mixed models assessed whether distress and pain each predicted the other in the following week. Results: 72 participants provided responses over 49 weeks. The participants had a median (IQR) age of 43 (37-51) years, 63% (n=45) were unemployed and most were females (n=51;71%). Distress and pain fluctuated concurrently within individuals: distress was positively correlated with worst pain ({rho}=0.66, 95% CI= 0.60-0.72, p<0.001) and average pain ({rho}=0.70, 95% CI=0.64-0.75, p<0.001) intensity within the same week. Worst pain (OR=1.42, 95% CI=1.17-1.71, p<0.001) and average pain (OR=1.43, 95% CI=1.20-1.71, p<0.001) intensity both predicted distress in the next week. Distress predicted worst pain intensity (OR=1.25, 95% CI=1.07-1.46, p=0.023) but not average pain intensity (OR=1.19, 95% CI=1.01-1.40, p=0.152) in the next week. Conclusions: The temporal relationship between distress and worst pain intensity was bidirectional, whereas distress did not temporally predict average pain intensity. Both pain and emotional distress should receive attention from HIV research and clinical care in low-resource settings.